47 research outputs found

    Computational identification of transposable elements in the mouse genome

    Get PDF
    Repeat sequences cover about 39 percent of the mouse genome and completion of sequencing of the mouse genome [1] has enabled extensive research on the role of repeat sequences in mammalian genomics. This research covers the identification of Transposable elements (TEs) within the mouse transcriptome, based on available sequence information on mouse cDNAs (complementary DNAs) from GenBank [28]. The transcripts are screened for repeats using RepeatMasker [23], whose results are sieved to retain only Interspersed repeats (IRS). Using various bioinformatics software tools as well as tailor made programming, the research establishes: (i) the absolute location coordinates of the TEs on the transcript. (ii) The location of the IRs with respect to the 5’UTR, CDS and 3’UTR sequence features. (iii) The quality of alignment of the TE’s consensus sequence on the transcripts where they exist, (iv) the frequencies and distributions of the TEs on the cDNAs, (v) descriptions of the types and roles of transcripts containing TEs. This information has been collated and stored in a relational database (MTEDB) at http://warta.bio.psu.edu/htt_doc/M TEDB/homepage.htm)

    Transduplication resulted in the incorporation of two protein-coding sequences into the Turmoil-1 transposable element of C. elegans

    Get PDF
    Transposable elements may acquire unrelated gene fragments into their sequences in a process called transduplication. Transduplication of protein-coding genes is common in plants, but is unknown of in animals. Here, we report that the Turmoil-1 transposable element in C. elegans has incorporated two protein-coding sequences into its inverted terminal repeat (ITR) sequences. The ITRs of Turmoil-1 contain a conserved RNA recognition motif (RRM) that originated from the rsp- 2 gene and a fragment from the protein-coding region of the cpg-3 gene. We further report that an open reading frame specific to C. elegans may have been created as a result of a Turmoil-1 insertion. Mutations at the 5' splice site of this open reading frame may have reactivated the transduplicated RRM moti

    U12 intron positions are more strongly conserved between animals and plants than U2 intron positions

    Get PDF
    We report that the positions of minor, U12 introns are conserved in orthologous genes from human and Arabidopsis to an even greater extent than the positions of the major, U2 introns. The U12 introns, especially, conserved ones are concentrated in 5'-portions of plant and animal genes, where the U12 to U2 conversions occurs preferentially in the 3'-portions of genes. These results are compatible with the hypothesis that the high level of conservation of U12 intron positions and their persistence in genomes despite the unidirectional U12 to U2 conversion are explained by the role of the slowly excised U12 introns in down-regulation of gene expression

    Retrophylogenomics place tarsiers on the evolutionary branch of anthropoids

    Get PDF
    One of the most disputed issues in primate evolution and thus of our own primate roots, is the phylogenetic position of the Southeast Asian tarsier. While much molecular data indicate a basal place in the primate tree shared with strepsirrhines (prosimian monophyly hypothesis), data also exist supporting either an earlier divergence in primates (tarsier-first hypothesis) or a close relationship with anthropoid primates (Haplorrhini hypothesis). The use of retroposon insertions embedded in the Tarsius genome afforded us the unique opportunity to directly test all three hypotheses via three pairwise genome alignments. From millions of retroposons, we found 104 perfect orthologous insertions in both tarsiers and anthropoids to the exclusion of strepsirrhines, providing conflict-free evidence for the Haplorrhini hypothesis, and none supporting either of the other two positions. Thus, tarsiers are clearly the sister group to anthropoids in the clade Haplorrhini

    RAB5A and TRAPPC6B are novel targets for Shiga toxin 2a inactivation in kidney epithelial cells

    Get PDF
    The cardinal virulence factor of human-pathogenic enterohaemorrhagic Escherichia coli (EHEC) is Shiga toxin (Stx), which causes severe extraintestinal complications including kidney failure by damaging renal endothelial cells. In EHEC pathogenesis, the disturbance of the kidney epithelium by Stx becomes increasingly recognised, but how this exactly occurs is unknown. To explore this molecularly, we investigated the Stx receptor content and transcriptomic profile of two human renal epithelial cell lines: highly Stx-sensitive ACHN cells and largely Stx-insensitive Caki-2 cells. Though both lines exhibited the Stx receptor globotriaosylceramide, RNAseq revealed strikingly different transcriptomic responses to an Stx challenge. Using RNAi to silence factors involved in ACHN cells’ Stx response, the greatest protection occurred when silencing RAB5A and TRAPPC6B, two host factors that we newly link to Stx trafficking. Silencing these factors alongside YKT6 fully prevented the cytotoxic Stx effect. Overall, our approach reveals novel subcellular targets for potential therapies against Stx-mediated kidney failure.publishedVersio

    Integrative Annotation of 21,037 Human Genes Validated by Full-Length cDNA Clones

    Get PDF
    The human genome sequence defines our inherent biological potential; the realization of the biology encoded therein requires knowledge of the function of each gene. Currently, our knowledge in this area is still limited. Several lines of investigation have been used to elucidate the structure and function of the genes in the human genome. Even so, gene prediction remains a difficult task, as the varieties of transcripts of a gene may vary to a great extent. We thus performed an exhaustive integrative characterization of 41,118 full-length cDNAs that capture the gene transcripts as complete functional cassettes, providing an unequivocal report of structural and functional diversity at the gene level. Our international collaboration has validated 21,037 human gene candidates by analysis of high-quality full-length cDNA clones through curation using unified criteria. This led to the identification of 5,155 new gene candidates. It also manifested the most reliable way to control the quality of the cDNA clones. We have developed a human gene database, called the H-Invitational Database (H-InvDB; http://www.h-invitational.jp/). It provides the following: integrative annotation of human genes, description of gene structures, details of novel alternative splicing isoforms, non-protein-coding RNAs, functional domains, subcellular localizations, metabolic pathways, predictions of protein three-dimensional structure, mapping of known single nucleotide polymorphisms (SNPs), identification of polymorphic microsatellite repeats within human genes, and comparative results with mouse full-length cDNAs. The H-InvDB analysis has shown that up to 4% of the human genome sequence (National Center for Biotechnology Information build 34 assembly) may contain misassembled or missing regions. We found that 6.5% of the human gene candidates (1,377 loci) did not have a good protein-coding open reading frame, of which 296 loci are strong candidates for non-protein-coding RNA genes. In addition, among 72,027 uniquely mapped SNPs and insertions/deletions localized within human genes, 13,215 nonsynonymous SNPs, 315 nonsense SNPs, and 452 indels occurred in coding regions. Together with 25 polymorphic microsatellite repeats present in coding regions, they may alter protein structure, causing phenotypic effects or resulting in disease. The H-InvDB platform represents a substantial contribution to resources needed for the exploration of human biology and pathology

    Mobilome of Apicomplexa Parasites

    No full text
    Transposable elements (TEs) are mobile genetic elements found in the majority of eukaryotic genomes. Genomic studies of protozoan parasites from the phylum Apicomplexa have only reported a handful of TEs in some species and a complete absence in others. Here, we studied sixty-four Apicomplexa genomes available in public databases, using a ‘de novo’ approach to build candidate TE models and multiple strategies from known TE sequence databases, pattern recognition of TEs, and protein domain databases, to identify possible TEs. We offer an insight into the distribution and the type of TEs that are present in these genomes, aiming to shed some light on the process of gains and losses of TEs in this phylum. We found that TEs comprise a very small portion in these genomes compared to other organisms, and in many cases, there are no apparent traces of TEs. We were able to build and classify 151 models from the TE consensus sequences obtained with RepeatModeler, 96 LTR TEs with LTRpred, and 44 LINE TEs with MGEScan. We found LTR Gypsy-like TEs in Eimeria, Gregarines, Haemoproteus, and Plasmodium genera. Additionally, we described LINE-like TEs in some species from the genera Babesia and Theileria. Finally, we confirmed the absence of TEs in the genus Cryptosporidium. Interestingly, Apicomplexa seem to be devoid of Class II transposons

    Lack of population differentiation patterns of previously identified putatively adaptive transposable element insertions at microgeographic scales

    No full text
    Background. Transposable elements (TEs) play an important role in genome function and evolution. It has been shown that TEs are a considerable source of adaptive changes in the genome of Drosophila melanogaster. Specifically, footprints of selection at the DNA level, the presence of population differentiation patterns across environmental gradients, and detailed mechanistic and fitness analyses of a few candidate adaptive TEs pointed to the role of TEs in environmental adaptation. However, whether the population differentiation patterns observed at large geographic scales can be replicated at a microgeographic scale has never been assessed before./nResults. In this work, we explored the population patterns of putatively adaptive TEs at a micro-spatial scale level. We compared the frequencies of TEs, previously identified as putatively adaptive and putatively neutral, in populations collected in opposite slopes of the Evolution Canyon at Mt. Carmel in Israel separated by 200 m on average. However, the differentiation patterns previously observed across large geographic distances (2000–2200 km) were not replicated at the microscale level of the Evolution Canyon populations./nConclusion. TE insertions previously associated with D. melanogaster environmental adaptation at a macro scale level do not play such a role at the microscale level of the Evolution Canyon populations. However, these results do not exclude a role of TEs in microgeographic adaptation because the dataset analyzed in this work is restricted to TEs identified in a single North American strain and as such is highly biased and incomplete.This work was supported by grants from the European Commission (PCIG-GA-2011-293860) and from the Spanish Government (BFU-2011-24397) awarded to JG and by the Institute of Bioinformatics funds awarded to WM. We acknowledge support of the publication fee by the CSIC Open Access Publication Support Initiative through its Unit of Information Resources for Research (URICI)

    Mammalian Overlapping Genes: The Comparative Perspective

    No full text
    It is believed that 3.2 billion bp of the human genome harbor ∼35,000 protein-coding genes. On average, one could expect one gene per 300,000 nucleotides (nt). Although the distribution of the genes in the human genome is not random,it is rather surprising that a large number of genes overlap in the mammalian genomes. Thousands of overlapping genes were recently identified in the human and mouse genomes. However,the origin and evolution of overlapping genes are still unknown. We identified 1316 pairs of overlapping genes in humans and mice and studied their evolutionary patterns. It appears that these genes do not demonstrate greater than usual conservation. Studies of the gene structure and overlap pattern showed that only a small fraction of analyzed genes preserved exactly the same pattern in both organisms
    corecore